Papers with domain filter
Parallel Corpus Filtering via Pre-trained Language Models (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods to filter out noisy parallel sentences from web crawled data are in demand. |
| Approach: | They propose a method to filter out noisy sentence pairs from web crawled corpora using pre-trained language models. |
| Outcome: | The proposed method outperforms baselines and achieves state-of-the-art on two datasets. |